Explore recent issues of Contract Pharma covering key industry trends.
Read the full digital version of our magazine online.
Behind every facility expansion, technology investment, and quality milestone in the CDMO sector is a leadership team making deliberate choices about where to focus, how to grow, and when to take calculated risks.
Stay informed! Subscribe to Contract Pharma for industry news and analysis.
Get the latest updates and breaking news from the pharmaceutical and biopharmaceutical industry.
Discover the newest partnerships and collaborations within the pharma sector.
Keep track of key executive moves and promotions in the pharma and biopharma industry.
Updates on the latest clinical trials and regulatory filings.
Stay informed with the latest financial reports and updates in the pharma industry.
A video roundup of the week’s top industry news stories.
Expert Q&A sessions addressing crucial topics in the pharmaceutical and biopharmaceutical world.
In-depth articles and features covering critical industry developments.
Access exclusive industry insights, interviews, and in-depth analysis.
Insights and analysis from industry experts on current pharma issues.
A one-on-one video interview between our editorial teams and industry leaders.
Listen to expert discussions and interviews in pharma and biopharma.
Contract Pharma Stream offers a centralized destination where users can watch expert-led sessions anytime, anywhere
A detailed look at the leading US players in the global pharmaceutical and BioPharmaceutical industry.
Browse companies involved in pharmaceutical manufacturing and services.
Comprehensive company profiles featuring overviews, key statistics, services, and contact details.
A comprehensive glossary of terms used in the pharmaceutical and biopharmaceutical industry.
Watch in-depth videos featuring industry insights and developments.
Download in-depth eBooks covering various aspects of the pharma industry.
Access detailed whitepapers offering analysis on industry topics.
View and download brochures from companies in the pharmaceutical sector.
Explore content sponsored by industry leaders, providing valuable insights.
Stay updated with the latest press releases from pharma and biopharma companies.
Explore top companies showcasing innovative pharma solutions.
Meet the leaders driving innovation and collaboration.
Engage with sessions and panels on pharma’s key trends.
Hear from experts shaping the pharmaceutical industry.
Join online webinars discussing critical industry topics and trends.
A comprehensive calendar of key industry events around the globe.
Live coverage and updates from major pharma and biopharma shows.
Find advertising opportunities to reach your target audience with Contract Pharma.
Review the editorial standards and guidelines for content published on our site.
Understand how Contract Pharma handles your personal data.
View the terms and conditions for using the Contract Pharma website.
What are you searching for?
Can the price of pharmaceutical molecules be predicted by artificial intelligence alone from molecular structure?
August 13, 2026
By: Michele Jermini
By: Paul Hanselmann
By: Federico Brooks
By: Enrico Polastro
Editor’s Take: AI can extract a meaningful pricing signal from molecular structure, but commercial factors still account for most of the variation.
Recent advances in cheminformatics and artificial intelligence (AI) have enabled increasingly accurate predictions of molecular properties, synthetic accessibility, and feasible synthetic routes directly from molecular structure.
Whether similar approaches can be extended to predict the commercial value of organic intermediates remains an open question.
To investigate this hypothesis, we have developed and trained a multi-layer machine-learning architecture using more than 7,000 pharmaceutical intermediates represented through SMILES strings and their market price.
The objective was not just to build a pricing engine but more fundamentally to evaluate how much pricing information is encoded in molecular structure and therefore extractable through modern AI techniques.
The model recovers approximately one-third of the observed price variability from just the molecular structure, demonstrating that molecular structure contains a measurable—but incomplete—economic signal that can be extracted using the machine-learning methods applied in this study.
These results highlight the potential but also the limitations of structure-based pricing models.
Over the last two decades, breakthroughs have been achieved in applying cheminformatics and artificial intelligence tools to molecular design and process development.
Examples include computer-assisted synthesis planning platforms like ASKCOS or Synthia, and molecular scoring systems such as Reaxys, SAS, SCScore and FSC. These tools estimate synthetic accessibility, route complexity and the synthesis feasibility of target molecules based on their structure.
While such approaches provide valuable information on how a molecule can be synthesized, these do not directly address the key question: Can the price of an organic intermediate be inferred from its molecular structure?
Prima facie, the hypothesis seems plausible because molecular structure influences synthesis complexity, stereochemical requirements, functional group density, precursor availability, and ultimately production costs.
However, contrary to melting point, toxicity, or solubility—price is not an intrinsic molecular property. Rather, it results from the interaction between chemistry, technology, supply chains, and competitive dynamics.
The objective of our work has therefore been to investigate the extent to which price information is encoded in molecular structure and whether modern AI methods can successfully extract and quantify this signal.
Tools like ASKCOS—an open-source software suite for computer-aided synthesis planning developed by MIT and E. Merck’s Chematica, and now called Synthia—allow chemists to rapidly identify feasible synthetic routes and their complexity to access target molecules whose structure has been translated into computer-readable form, such as SMILES strings or Molfiles.
Similarly, multiple systems like SAS (Synthesis Accessibility Score), SCS (Synthesis Complexity Score), or FSC (Focused Synthesizability Score) based on Machine Learning trained on millions of different reactions are routinely used in early-stage drug discovery to weed out molecules with high theoretical scores, a proxy for difficult, if not impossible, synthesis.
However, while these tools are useful in assessing how a molecule can be synthesized by providing a quantitative view of its synthesis complexity, they do not say what it will cost—synthesis complexity and price being related but not identical concepts.
Machine-learning approaches for predicting the price of chemical compounds have so far followed three main directions.
The first relies on graph neural networks trained directly on large commercial catalogs, learning the relationship between molecular structure and listed prices.1 While this approach has the advantage of covering very large datasets—the representativeness of catalog prices is liable to be challenged. Consequently, these models learn a combination of structural and commercial effects rather than the contribution of molecular structure alone.
A second approach estimates price from predicted synthetic routes combined with the cost of starting materials.2 These methods can potentially achieve higher accuracy when reliable retrosynthetic information is available – but are dependent on the quality of the route prediction itself and cannot be considered structure-only approaches
A third direction derives accessibility or cost-related scores from market data using contrastive or self-supervised learning techniques.3 These methods are mainly intended to rank compounds according to expected cost or accessibility rather than predict an absolute market price
A molecule can be synthetically challenging yet inexpensive if produced at a large scale through an optimized process. Conversely, a structurally simple molecule may have a high price due to puny volumes, scarce feedstock availability, or stringent containment requirements due to its potency.
The development of AI models trained on large product datasets and market prices to estimate the price of organic intermediates directly from molecular structure would therefore represent a compelling innovation and address substantial unmet needs.
Such models extending the structure-based analysis beyond synthetic feasibility, combining it with a price prediction, would have the potential to provide early in the product development process data-driven guidance on cost determinants, allowing screening for alternative intermediates or synthesis pathways. Expected benefits would include supporting informed decision-making across R&D and procurement, providing an integrated, predictive tool that bridges chemistry and economics.
To this end, we have combined the multi-year experience accumulated in the development and sourcing of organic intermediates for the pharmaceutical industry with expertise in AI and MLM (Machine Learning Models) for building a model to test the feasibility of predicting the price of new intermediates based solely on their molecular structure and assessing how much of the price can be traced to this single variable.
Dataset
The input data comprise more than seven thousand commercially relevant pharmaceutical intermediates. For each compound, the molecular structure was represented by its SMILES string together with the latest market price; the CAS number was used for compound identification where available.
To reduce variability associated with geography, exchange rates, and transaction scale, all prices were normalized using quotations from Chinese producers for volumes of 1,000 kg and applying 2026 RMB/USD exchange rates.
The key descriptors of the molecule are provided by CAS/SMILES in a machine-readable format “understandable” by neural network models. At the same time, the price is the target variable that the model aims to learn and predict, linking it to the structural features of the molecule.
Although a dataset of this size is significant by industrial standards, it represents only a small fraction of the broader chemical universe. Consequently, some prediction errors are inevitably attributable to an incomplete coverage of chemical space. However, as discussed later, several systematic behaviors observed in the model suggest that limited dataset size alone does not fully explain the observed performance limits.
Model architecture
The architecture of the model developed is schematically illustrated in Figure 1.
Figure 1. AI-Based Molecular Price Prediction Model Architecture
It consists of a multilayered machine learning system designed to estimate the price of organic pharmaceutical intermediates in USD/kg directly from their molecular structure. It combines traditional cheminformatics, modern machine learning techniques, and large language model (LLM)–derived insights into a single predictive framework.
Features
The first step consists of converting the molecular structure into a set of numerical descriptors that machine-learning algorithms can process. These descriptors include well-established chemical characteristics such as molecular weight, polarity, ring structures, and other features reflecting the molecule composition and complexity.
In parallel, additional descriptors are generated using a Large Language Model (LLM) to capture higher-level chemical information, like perceived synthetic complexity and key functional groups.
As this process generates multiple variables, only those providing meaningful information are retained. The selected descriptors can also be combined to create new variables, helping to identify more complex relationships between molecular structure and price.
The dataset is then divided into separate subsets used for training, validation, and testing. To ensure that the results are robust and not dependent on a particular data split, the model is trained and evaluated several times using different data subsets, which also reduces overfitting risk.
The dataset was divided into independent training, validation, and test subsets.
Model development, parameter optimization, and architecture selection were performed exclusively on the training and validation data.
Final performance was evaluated on a held-out test set comprising molecules never seen during model training. Cross-validation procedures were additionally applied to reduce dependence on any specific train-test split and provide a more robust assessment of predictive performance.
The model combines several complementary artificial intelligence approaches, each designed to identify different relationships between molecular structure and price. Some models focus on simple and easily interpretable relationships, while others capture more complex patterns that may not be immediately apparent.
In addition, specialized models are trained to perform particularly well on specific families of molecules or regions of chemical space. This allows the system to adapt its behavior depending on the type of compound analyzed.
The predictions generated by the various models are subsequently combined into a single estimate. Additional refinements include comparing the molecule under consideration with structurally similar compounds in the training dataset—the hypothesis being that similar molecules often exhibit comparable pricing patterns.
A final AI-based review step is then applied to identify unusual cases and improve predictions for rare or particularly complex molecular structures.
Ultimately, the multilayered predictive system involving multiple models working together receives the molecular structure of a compound as input and generates the estimated market price as output.
Table 1 summarizes dataset size and predictive performance.
5.1 General predictive performance
The model’s ability to assess molecules that it had not been exposed to has been evaluated through an independent test set comprising 869 pharmaceutical intermediates.
The overall performance achieved an R² value of 0.36 on the price logarithm, corresponding to a Root Mean Square Error (RMSE) of 0.56 and a Mean Absolute Error (MAE) of 0.44 (see Figure 2).
Figure 2. Predicted vs. Actual Pharmaceutical Intermediate Prices in the Holdout Test Set
The R² value of 0.36 indicates that approximately one-third of the observed price variability can be recovered from just the molecular structure. The remaining variability derives from non-structural price determinants and/or structural information that the model does not fully capture.
These findings suggest that based solely on the molecular structure, the model can explain approximately one-third of the observed variation in market prices – confirming that molecular structure provides a measurable economic signal and that machine-learning techniques can partly link chemistry with commercial value.
However, the practical implications of these metrics require careful interpretation. The typical prediction error for an individual molecule remains approximately a factor of 2.3, with only 43% of compounds predicted within a factor of two of their actual market prices.
Therefore, while the model successfully distinguishes broad pricing regions within the chemical space, its ability to predict the absolute prices of specific products is limited, confirming that molecular structure alone provides incomplete information on pricing.
5.2 Systematic prediction compression
Analysis of the predicted-versus-observed prices reveals a clear systematic behavior. Predictions are compressed towards the center of the distribution, with a best-fit slope of approximately 0.32 compared with the ideal value of 1.0.
In practical terms, the model tends to overestimate the cheapest molecules and underestimate the most expensive ones:
• The lowest-priced 20% of compounds being predicted to have a price of approximately 64 USD/kg versus an average actual market price of around 16 USD/kg.
• The highest-priced 20% having an average market price in the order of 2,231/USD/kg to be compared to the 339 USD/kg predicted by the model.
The observed compression suggests that molecular structure captures broad determinants of cost but not the drivers responsible for extreme prices.
5.3 Performance after removal of price extremes
To evaluate the influence of extreme values, molecules priced below USD 20/kg and above USD 20,000/kg were removed from the test set.
The resulting improvement was limited—the median prediction error decreasing from approximately 2.3-fold to 2.1-fold, while the proportion of compounds predicted within a factor of two increased from 43% to 48%.
The modest improvement obtained after removing outliers suggests that the model’s limitations are not confined to a few exceptional compounds but are instead linked to broader structural characteristics.
5.4 What the model reveals about chemical price formation
Taken together, these results support two important conclusions.
First, molecular structure undeniably contains information relevant to pricing. Functional groups, stereochemistry, molecular complexity and other structural features influence synthetic accessibility and manufacturing costs. The model successfully extracts part of this information, using it to explain approximately one-third of the observed price variation.
Second, molecular structure alone does not explain most market price differences.
Unlike physicochemical properties such as molecular weight, boiling point, density, solubility, or toxicity, price is not an intrinsic molecular property. Rather, it results from the interaction between molecular structure, manufacturing technology, production scale, supply-chain structure, and competitive dynamics. This distinction helps explain why predicting price from structure alone is inherently more difficult than predicting traditional molecular properties.
Our study suggests that about one-third of the variability in pharmaceutical intermediate prices can be inferred directly from the molecular structure using the machine-learning techniques applied.
Although the resulting predictions are insufficient for accurate commercial pricing, these confirm that molecular structure contains a measurable economic signal.
Substantial improvement of predictive performance can be expected using future models combining structural descriptors with synthetic processes and commercial information.
References
1. Sanchez-Garcia, R.; Havasi, D.; Takács, G.; Robinson, M. C.; Lee, A.; von Delft, F.; Deane, C. M. CoPriNet: graph neural networks provide accurate and rapid compound price prediction for molecule prioritization. Digital Discovery 2023, 2, 103–111. doi:10.1039/D2DD00071G.
2. Abderrahmane, M.; Tajmouati, H.; Barros Ribeiro da Silva, V.; Perron, Q. Predicting the price of molecules using their predicted synthetic pathways. Molecular Informatics 2025, 44 (2), 202400039. doi:10.1002/minf.202400039.
3. Hastedt, F.; Hellgardt, K.; Yaliraki, S.; et al. MolPrice: assessing synthetic accessibility of molecules based on market value. Journal of Cheminformatics 2025, 17, 150. doi:10.1186/s13321-025-01076-3.
Dr. Michele Jermini is the Managing Director of Exeris ([email protected]).
Dr. Paul Hanselmann is Founder and CEO of ChemSynthDesign GmbH.
Federico Brooks is Founder and CTO of Custodian Privacy Sagl.
Dr. Enrico Polastro is a Vice-President of Arthur D.Little.
Enter your account email.
A verification code was sent to your email, Enter the 6-digit code sent to your mail.
Didn't get the code? Check your spam folder or resend code
Set a new password for signing in and accessing your data.
Your Password has been Updated !